Back

Computational Psychiatry

Ubiquity Press, Ltd.

Preprints posted in the last 90 days, ranked by how well they match Computational Psychiatry's content profile, based on 12 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
Mapping abstraction and metacognition onto distinct transdiagnostic symptom profiles

Oka, T.; Kunisato, Y.; Koizumi, K.; Murakami, M.; Six, H.; Taylor, J. E.; Cortese, A.

2026-06-22 psychiatry and clinical psychology 10.64898/2026.06.17.26354096 medRxiv
Top 0.1%
22.0%
Show abstract

Transdiagnostic psychiatric research on reward-guided learning has largely focused on simple associative processes, leaving it unclear whether or how higher-level processes are disrupted. Here, we studied how abstraction, the ability to extract relevant features from complex information, and metacognition, the ability to monitor and evaluate one's own mental processes, map onto specific transdiagnostic dimensions. Using an online sample (N = 249), we examined associations between these processes and three cross-culturally robust transdiagnostic dimensions derived from a large existing dataset (N = 19,505): Compulsive hypersensitivity, Social withdrawal, and Addictive behaviours. Computational modelling of an abstract representation learning task with confidence judgments revealed that Compulsive hypersensitivity was negatively associated with both abstraction ability (pboot = 0.003) and metacognitive sensitivity (pboot = 0.005), while Social withdrawal was positively associated with metacognitive sensitivity alone (pboot = 0.002). Moreover, transdiagnostic dimensions revealed more coherent associations with higher-order cognition than symptom-level analyses, highlighting the added value of examining psychopathology at the factor rather than the symptom level. These findings portray a hierarchical view of cognitive dysfunctions in psychopathology and point to representational and metacognitive processes as potential targets for transdiagnostic intervention.

2
Deciding when to decide: How recency, urgency, risk, and bias shape human sequential decision-making: A case study across the obsessive-compulsive spectrum

Abdelrazik, A. H.; Dayan, P.

2026-08-25 neuroscience 10.64898/2026.08.21.746184 medRxiv
Top 0.1%
10.9%
Show abstract

Deciding when to stop gathering information and commit to a choice is a fundamental challenge in decision-making under uncertainty. Normative characterizations such as Partially Observable Markov Decision Processes (POMDPs) prescribe mathematically optimal stopping rules; however, human evidence gathering systematically departs from optimality. Pathological departures -- such as the excessive indecisiveness characteristic of obsessive-compulsive disorder (OCD) -- offer an important opportunity to investigate the cognitive mechanisms involved in stopping. We extend a POMDP framework to incorporate key candidate suboptimalities: a biased prior belief, transient evidence exaggeration, progressive forgetting, boosted costs of error, temporal regulation (patience and urgency), and misperception of a deadline. We evaluate this model in a pre-existing dataset comprising 105 participants spanning healthy controls, generalised anxiety disorder, and the OCD spectrum performing an information gathering task with controlled, stochastic, deadlines. Model comparison reveals that human sequential choices are broadly governed by subjective risk penalties and time-dependent urgency, with a smaller and less certain contribution from an over-weighting of recent evidence, which a random-effects comparison does not support at the population level. Individuals differ in how that over-weighting is implemented: in one deadline condition, subjects divide almost evenly between models carrying a transient exaggeration of the newest sample, models carrying progressive forgetting of older evidence, and models carrying no recency mechanism at all. Crucially, while risk sensitivity and choice stochasticity act as shared mechanisms across conditions, mechanisms such as belief bias and patience are more variable. Finally, using OCD as a clinical case study, we demonstrate that simulating choices from the fitted exaggeration model reproduces model-agnostic regression signatures of clinical indecision, which the forgetting and no-recency accounts do not. These findings offer a generative foundation for dissecting clinical departures in information gathering across the obsessive-compulsive spectrum.

3
Two modes of aversive control in suicidality: joint computational modelling exposes regime-specific clinical signatures invisible to symptom-based stratification

Laessing, P.; Karvelis, P.; Rashid-Cocker, A. S.; Ruocco, A. C.; Koudys, J. W.; Kennedy, J. L.; Zai, C. C.; Dayan, P.; Diaconescu, A.

2026-06-11 psychiatry and clinical psychology 10.64898/2026.06.09.26355278 medRxiv
Top 0.1%
7.1%
Show abstract

Suicidal thoughts and behaviours (STBs) are heterogeneous in their proximal dynamics, planning, and stress-sensitivity, yet most subtyping efforts remain symptom-driven and rarely validated across independent datasets. Computational mixture modelling offers a principled alternative: by fitting explicit models of learning and action selection and partitioning individuals by their latent parameter profiles, it can identify mechanistically distinct control strategies invisible to cross-sectional symptom measurement. We applied this approach to aversive Go/NoGo performance, jointly clustering two independently collected STB-enriched samples (N = 50 and N = 184) using tasks with the same structure but different duration, reversal timing, and clinical instrumentation. Two recurrent behavioural regimes emerged: a fast/adaptive regime characterised by rapid policy updating and elevated feedback reactivity, and a slow/perseverative regime characterised by slow updating, high choice determinism, and a pronounced cost following contingency reversal. These regimes were stable across initialisations, recovered more parsimoniously in joint than independent solutions, and were largely orthogonal to symptom-based stratification. Critically, stratification by regime exposed clinical-computational coupling structures substantially attenuated in pooled analyses. Pooled, population-level associations were modest and anchored by a broad affective burden axis. Within the slow/perseverative regime, coupling reorganised around learning dynamics and internalizing burden (depression, hopelessness, and active suicidal ideation) with markedly larger effect sizes. Within the fast/adaptive regime, a dissociation between anxious-compulsive and antisocial-disinhibitory profiles emerged along the same computational axis, invisible at the population level. These findings support a view of suicidality heterogeneity in which clinically similar individuals differ in the control strategies they recruit under aversive uncertainty - variation that symptom measurement alone cannot capture.

4
Modelling Metacognition: A Joint Prediction-Confidence Model for Predictive Inference Task Data

Martinez, E. F.; Waade, P. T.; Heinzle, J.; Hess, A. J.

2026-08-21 neuroscience 10.64898/2026.08.14.744790 medRxiv
Top 0.1%
6.5%
Show abstract

Metacognition is the ability to reflect on and evaluate our own cognitive processes. It is often altered in psychopathology. Yet, the computational mechanisms underlying these alterations remain unclear. In this work, we extend Hierarchical Gaussian Filter (HGF) models to jointly fit trial-by-trial predictions and confidence ratings in a predictive inference task, providing an individualised characterisation on metacognitive processing. Applying our cognitive computational model to a large subclinical open dataset (N=430), we are able to achieve, on average, excellent fit of prediction responses [Formula] and a moderate to good fit of confidence ratings [Formula]. Analysis of experimental change-points revealed that our model accurately captures confidence self-reports dynamics around these change-points. Posterior parameter estimates reveal a negative effect of sensory input prediction errors and a positive effect of sensory input prediction precision on confidence ratings, respectively. In addition, we replicate state-of-the-art findings related to compulsivity as measured by a transdiagnostic factor score, such as inflated confidence and a decoupling of action updates (here, prediction errors) and confidence in compulsivity. These results demonstrate the robustness of our methodology and the potential of joint prediction-confidence modelling to uncover latent metacognitive alterations in psychopathology.

5
Gambling disorder symptom severity and mental health: an item-response-theory analysis of DSM-5 criteria for gambling disorder

Vogl, F.; Wolff, H.-G.; Buth, S.; Peters, J.

2026-08-13 psychiatry and clinical psychology 10.64898/2026.08.12.26360267 medRxiv
Top 0.1%
5.5%
Show abstract

The present study examined the relationship between the DSM-5 diagnostic criteria for gambling disorder (GD) and gambling severity via an item response theory (IRT) analysis in two large German population survey data sets (Buth et al. (2022, 2024)). IRT-based person fit analyses may reveal atypical response patterns (e.g. endorsing criteria linked to higher levels of disorder severity, but not criteria linked to lower levels). We examined the link of such atypical response patterns and mental health as measured by the MHI-5, employing a 2-parameter-logistic (2PL) IRT model and a linear mixed model with random intercepts. Results largely replicated previously reported item severity rankings across both samples: GD criteria such as loss chasing and a preoccupation with gambling were generally linked to lower severity levels, whereas criteria such as withdrawal symptoms or job/family problems where generally linked to higher severity levels. Modelling revealed a reduced assessment sensitivity in lower gambling severity ranges. Furthermore, person fit analyses suggest that atypical symptom patterns may be linked to poorer mental health (MHI-5). Implications for the interpretability of total scores of endorsed criteria and the validity of diagnostic practices determining eligibility for treatment and financial compensation are discussed.

6
The relationship between serotonin transporter occupancy and extracellular serotonin concentration is hyperbolic, not linear: implications for safely tapering antidepressants

Cohrs, D.; Shapiro, B.

2026-06-18 psychiatry and clinical psychology 10.64898/2026.06.09.26355019 medRxiv
Top 0.1%
5.2%
Show abstract

Background: Hyperbolic tapering is an increasingly recognized approach for discontinuing serotonin reuptake inhibitor (SRI) antidepressants that involves non-linear dose reductions with equal stepwise reductions in serotonin transporter (SERT) occupancy to mitigate withdrawal symptoms. Its theoretical basis is the hyperbolic relationship between SRI dose and SERT occupancy reported in radioligand imaging studies. Hyperbolic tapering implicitly assumes that changes in SERT occupancy approximate changes in biologic effect and withdrawal risk. Because SERT occupancy plateaus across the therapeutic dose range of SRIs, this framework predicts relatively small biologic effects and withdrawal risk within this range. However, SERT occupancy influences serotonergic activity only indirectly via its effects on extracellular serotonin concentrations, and the relationship between these two variables is poorly characterized. Methods: We developed a two-pathway clearance model derived from mass-action kinetics to evaluate the steady-state relationship between SERT occupancy and extracellular serotonin concentrations under chronic SRI treatment. Results: Our analysis indicates that serotonin concentrations increase hyperbolically as transporter occupancy increases, suggesting that biologically meaningful differences in serotonergic signaling persist across the therapeutic dose range of SRIs despite plateauing occupancy. Conclusions: Our model predicts a hyperbolic relationship between SERT occupancy and extracellular serotonin concentrations, suggesting that changes in occupancy may not map proportionally onto serotonergic effect. These findings provide a potential mechanistic explanation for dose-dependent clinical effects of SRIs despite plateauing transporter occupancy and generate testable hypotheses regarding antidepressant tapering strategies. Empirical validation is warranted.

7
Task-Based Value Generalization Correlates With Positive Overgeneralization and Bipolar Symptoms

Li, J.; Malaviya, M.; Bennett, D.; Radulescu, A.

2026-07-17 neuroscience 10.64898/2026.07.12.737635 medRxiv
Top 0.1%
4.7%
Show abstract

Positive overgeneralization - the tendency to generalize from specific successes to broad expectations of future reward - has been linked to vulnerability to mania. Because positive overgeneralization has primarily been assessed using self-report measures, we have limited insight into the underlying cognitive process. Here, we introduce a behavioral paradigm designed to quantify how learned value generalizes to novel stimuli. We quantify individual generalization profiles by fitting psychometric functions to choice data. In an online transdiagnostic study (N=163), we show that task-based breadth of reward generalization is associated with both higher self-reported positive overgeneralization and subclinical bipolar symptoms. To provide a computational account of positive overgeneralization, we implement a reinforcement-learning model in which self-efficacy modulates the influence of anticipated future value during learning. We show that increasing this modulation reproduces the broader value propagation observed empirically. Together, these findings provide a behavioral and computational framework for studying positive overgeneralization, and suggest a mechanistic pathway by which success-related shifts in value representations may bias learning in ways relevant to bipolar risk.

8
Attenuation of value-to-evidence translation drives biased decision making in anxiety and depression

Gopnarayan, M. N.; Sheng, F.; Platt, M. L.; Ramakrishnan, A.

2026-06-11 neuroscience 10.64898/2026.06.09.731091 medRxiv
Top 0.1%
4.2%
Show abstract

Anxiety and depression are globally prevalent conditions associated with maladaptive decision making. However, whether affective symptoms primarily amplify threat avoidance or dampen motivational drive remains debated, and behavioural studies yield inconsistent findings. Here we show that both anxiety and depression impair the fundamental cognitive process of translating objective value into decision evidence. Across independent cohorts from the US and India, participants evaluated risky gambles while we assessed choice behaviour and the centroparietal positivity, an EEG marker of accumulating decision evidence. Prospect theory parameters, like risk and loss aversion, showed little association with symptom severity. Conversely, hierarchical drift-diffusion modelling revealed that higher symptom scores predicted attenuated value sensitivity during evidence accumulation, whereas decision caution remained intact. This reduction in value sensitivity suggests internalizing symptoms disrupt choice at the value-to-evidence interface, offering a unified mechanism underlying biased decision making in affective disorders.

9
Evidence-guided AI regularization for suicidal ideation prediction in pediatric bipolar disorder

Jabbar Abdl Sattar Hamoudi, H.; Wu, M.-J.; Sanches, M.; Zunta-Soares, G. B.; Soutullo, C. A.; Soares, J. C.; Mwangi, B.

2026-06-22 psychiatry and clinical psychology 10.64898/2026.06.18.26355841 medRxiv
Top 0.1%
4.2%
Show abstract

Background: Suicide prediction models in psychiatry often rely on purely data-driven feature selection, which can produce unstable and clinically opaque predictor sets in modest-sized samples. We developed Evidence-Based AI LASSO (EBAL), an evidence-guided regularization framework that incorporates curated clinical evidence into feature-specific penalty factors for interpretable prediction. Methods: Baseline data from 136 youth with confirmed bipolar spectrum disorder in the Greater Houston Area Bipolar Registry were analyzed using 20 candidate clinical predictors. Forty higher-level evidence documents on suicidality and related predictor domains were curated through a structured evidence synthesis workflow and indexed as an auditable evidence corpus. An open-weight large language model assigned feature-specific penalty factors using a prespecified scoring rubric, and these penalties were used to fit a weighted LASSO model. EBAL was compared with a standard evidence-agnostic LASSO using nested leave-one-out cross-validation. Results: For suicidal ideation, EBAL achieved an AUROC of 0.768, balanced accuracy of 0.757, sensitivity of 0.758, and specificity of 0.757. The standard LASSO achieved an AUROC of 0.760 and balanced accuracy of 0.715. EBAL improved balanced accuracy (+0.042, p=0.010) and Matthews correlation coefficient (+0.079, p=0.010), while retaining fewer stable predictors than standard LASSO (11/20 vs 18/20). The strongest positive predictors were current depressed mood, duration of mood disorder illness, and comorbid generalized anxiety disorder. For suicidal behavior, both models performed near chance and retained all candidate predictors. Limitations: The study was cross-sectional, single-site, and modest in sample size, with no external validation cohort. Conclusions: EBAL produced a sparser and more clinically coherent model for suicidal ideation in pediatric bipolar disorder, but did not improve prediction of suicidal behavior. These findings support evidence-guided regularization as a transparent strategy for aligning psychiatric prediction models with prior clinical knowledge while preserving interpretability.

10
Optimal Practice Schedules in a Dual-Rate Model of Motor Adaptation, and Their Recovery by Reinforcement Learning

Jeter, R.; Todorov, D.; Molkov, Y.

2026-06-22 neuroscience 10.64898/2026.06.17.732970 medRxiv
Top 0.1%
4.0%
Show abstract

A clinician guiding a stroke patient through a 45-minute rehabilitation session, a coach planning a training day, a teacher choosing the order of practice problems, they all face the same question: "given everything practiced so far, what should the next trial be?" The motor-learning literature offers two coarse answers, blocked and interleaved ("random") practice, with a well-known dissociation, blocked practice gives faster acquisition but worse retention, while interleaved practice gives the opposite. We argue that this dissociation is not a fixed property of practice schedules but a shadow of a richer structure. In particular, for a learner whose memory has a fast shared component and slower context-specific components, the best schedule should be a function of the learners current internal state and the time remaining before the retention probe. We make this precise in a minimal two-context fast-slow learner model whose optimal schedules can be computed exactly for short sessions and approximated by a structured beam-search upper bound for longer ones. The optimal schedule is not blocked, not interleaved, and not a single rule; it is a family of schedules determined by how much retention is weighted relative to acquisition. The family has three regimes (alternating, mixed, blocked-with-late-correction) and for long sessions, the optimal schedule has an interpretable structure -- exploit one context, repair the neglected one, then interleave to lock in retention. We then investigate whether a reinforcement-learning teacher, observing only the learners actions and errors without access to their internal memory states, can learn these optimal policies from interaction alone. Comparing these learned policies against the exact optima, we show that a model-free agent (PPO) recovers the short-horizon schedules and the long-horizon block-repair-interleave motif in the intermediate regime, but the benchmark also exposes a sharp failure in the acquisition-dominated regime, where PPO collapses to pure blocking and misses a sparse terminal correction. A warm-start diagnostic shows this failure is a genuine metastability of policy gradients rather than a tuning artifact, with blocked-plus-switch and pure-blocked acting as competing attractors that PPO cannot stabilize between. A hyperparameter sweep over observation history reveals that the agent requires very little behavioral context to plan optimally, demonstrating that partial observability is not a major barrier to finding optimal practice schedules. Finally, we discuss the implications of our framework for motor adaptation and contextual interference, offering practical insights on how instructors can design finite practice sessions to favor long-term retention.

11
Recovery from catatonia follows a hierarchical precision order

Saito, H.; Takizawa, Y.; Tateno, A.; Theorell, J.; Arakawa, R.; Tiger, M.

2026-08-17 psychiatry and clinical psychology 10.64898/2026.08.13.26360194 medRxiv
Top 0.1%
3.2%
Show abstract

Catatonia offers no principled basis for sequencing interventions when first-line benzodiazepines fail. Electroconvulsive therapy (ECT) is the established next step, but specifies what to escalate to, not what to stabilise first. Here we show that recovery is a constrained progression across five precision domains of hierarchical inference: sensory ({pi}s), policy ({beta}), motivational ({pi}m), and fast and slow volatility precision ({pi}v_fast, {pi}v_slow). In twenty-five consecutive inpatients managed without ECT, an order {pi}s [-&gt;] {beta} [-&gt;] {pi}m [-&gt;] {pi}v_fast [-&gt;] {pi}v_slow) held without inversion in every patient. Bush-Francis Catatonia Rating Scale (BFCRS) scores fell from 26.3 {+/-} 5.6 to 2.0 {+/-} 2.4 (p < 0.001), and functional recovery tracked restoration of organized action rather than symptom suppression. The framework predicts, untested in this uniformly remitting cohort, that interventions effective at one stage may destabilise another. Within stated conditions, a single inversion falsifies the ordering.

12
Factor Analysing Predictive Processing: No Evidence for a General Factor Across Tasks

Miller-Silva, C.; Knolle, F.; Greve, A.; de Beer, F.; Mujirishvili, T.; MacGregor, L. J.; Corlett, P. R.; Haarsma, J.; Powers, A. R.; Murray, G. K.

2026-06-18 psychiatry and clinical psychology 10.64898/2026.06.09.26354804 medRxiv
Top 0.1%
3.2%
Show abstract

Background & Hypothesis: Dysfunctional predictive processing (PP), specifically the aberrant weighting of priors, is a frequently-proposed mechanism for psychosis and psychosis-like phenomena (schizotypy). Evidence for this theory mostly originates from single-task studies, which assume that all tasks load onto a single latent construct of PP performance, but the underlying factor structure of PP tasks is unknown. PP deficits in psychosis may be better described by a two-factor, hierarchical model: weakened lower-level (perceptual) priors compensated by higher-level (cognitive) priors. Study Design: This study implements a multi-paradigm approach in healthy participants to investigate latent constructs underlying PP and their relationship to schizotypy. Participants (N = 73) completed 6 tasks measuring reliance on priors across language, memory, visual, and auditory domains. A factor analysis investigated whether performance across tasks is captured by a single or two-factor model. Study Results: Although a two-factor model best described performance, factors reflected within-task correlations rather than a PP hierarchy. Cross-task PP measures were poorly correlated, suggesting that individuals' weighting of priors was task-specific. A full model including all task outcomes (not factors) significantly predicted the severity of schizotypal aberrant beliefs but no other schizotypal measures. Conclusions: These results do not evidence a single factor underpinning PP performance. It is therefore inappropriate to use results from single tasks to propose a generalised PP deficit in psychosis. Variation was also not captured by a two-factor hierarchical model of priors. Further multi-paradigm research is required to evaluate alternative models or additional variables that describe aberrant PP in psychosis.

13
Detection of Frustration-related Operant Behavior in Rats via Machine Learning Methods

Wang, J.; Babu, A. S.; Nguyen, B.; Contreras, Y. M.; Shah, P.; Ramirez, I. C.; Green, T. A.

2026-09-01 animal behavior and cognition 10.64898/2026.08.26.747319 medRxiv
Top 0.1%
2.7%
Show abstract

Despite its strong link to neuropsychiatric conditions, frustration remains critically understudied in humans and animals alike. Therefore, there is an urgent need to develop tools to understand and therapeutically target frustration-related functions. Interestingly, humans and rats respond similarly during frustrative nonreward by increasing barpress durations. We previously validated barpress duration in rat operant tasks as a reliable measure of frustration-related behavior; however, it is wellknown that in addition to duration of responding, emotional states such as frustration alter other aspects of responding such as force of pressing. One-dimensional, static measures such as maximum force could miss rich information contained within operant data. Thus, the objective of this study is to apply machine learning (ML) to force/time profiles to discriminate frustration-related barpresses from non-frustration-related barpresses. Results showed an AUROC for FR1 (i.e., non-frustrated) vs. extinction (frustrated condition) for individual barpresses of 0.65 that improved to 0.84 with a chunk size of 10. The model generalized well to progressive ratio responding, a different kind of frustration procedure. We conclude that force/time profiling does provide utility beyond one dimensional measures of duration or force separately, meaning that we can indeed infer the internal state of frustration from behavior using ML techniques. Importantly, this project will also serve as proof-of-concept for applying ML to predict other internal states from barpress data.

14
Conversational trajectory degrades large language model detection of suicidal ideation relative to clinicians: a preregistered study

Kalinich, M.; Luccarelli, J.; Santa Maria, J.; Flathers, M.; Nguyen, A.; Song, S. H.; Makhoul, K.; Rivera Criado, M. J.; Ginapp, C. M.; Hill, B.; Shumate, J. N.; Notsu, H.; Smith, C.; Moss, F.; Torous, J.

2026-07-14 psychiatry and clinical psychology 10.64898/2026.07.10.26357132 medRxiv
Top 0.1%
2.6%
Show abstract

Background General-purpose large language models increasingly encounter emotional and therapy-like conversation, yet are not developed or evaluated as clinical systems. Existing safety evaluations rely largely on brief exchanges, although harms often unfold over extended interactions. Whether models maintain safety-relevant performance as conversations accumulate context remains unknown. Methods In this preregistered study, 400 clinician-validated statements, with or without suicidal ideation, were inserted at 0-200 speaker turns in 5 psychotherapy and 3 synthetic transcripts. Forty-nine LLMs and 8 clinicians performed the same binary classification task. Mixed-effects models estimated the effects of conversational depth, model scale, and model version on F1. Twelve top models were tested to 1,500 turns across conversational trajectories, with or without instruction restatement. Results F1 declined with depth across model families (p<0.001). Larger, newer models performed better but still degraded. Clinicians showed no decline (mean F1 0.86 at both 0 and 200 turns), but eight of nine proprietary models exceeded their performance at 200 turns. Conversational content, not length alone, explained F1 changes; the largest decrease was under adversarial context (p<0.001). Restating instructions increased F1 on therapy to near baseline (median {Delta}F1 +0.12; p<0.001; 89% median recovery) versus MSJ ({Delta}F1 +0.08; p=0.04; 38% recovery). Conclusions LLM detection of suicidal ideation degraded with conversational depth and trajectory, whereas clinician performance remained stable despite the strongest models exceeding most clinicians in absolute performance. Mental health AI safety evaluations should test sustained performance across realistic and adversarial trajectories rather than relying on short-prompt benchmarks.

15
Cross-Modal Benchmarking of Acoustic Prosody and Ventral Striatal BOLD for Depression-Related Anhedonia Classification: A Pre-Registered Study with the ClinicalWhisper Pipeline

Zhou, C.; Wu, M.; Xiang, Y.; Itti, L.

2026-06-11 neuroscience 10.64898/2026.06.08.728970 medRxiv
Top 0.1%
2.5%
Show abstract

Can computational analysis of a brief voice recording classify depression-related anhedonia as effectively as task-based fMRI? We address this question through a pre-registered cross-modal benchmarking study (osf.io/bsvrj) that evaluates two independent classification pipelines against depression-related anhedonia operationalized via self-report. Anhedonia-- the diminished capacity to experience pleasure or motivation to pursue rewards--is a transdiagnostic marker of reward-system dysfunction that predicts treatment resistance in major depressive disorder, yet current assessment requires either expensive functional neuroimaging or subjective self-report scales, neither of which scales to routine screening. We benchmarked two independent classification pipelines: Stream A extracted 88 acoustic prosody features from DAIC-WOZ clinical interviews (n = 142) using the ClinicalWhisper pipeline (Whisper Large-v3, pyannote diarization, OpenSMILE eGeMAPS v02); Stream B extracted nucleus accumbens BOLD activation during a reward task from the UCLA ds000030 dataset (n = 272). Three classifiers (logistic regression, random forest, gradient-boosted trees) were evaluated under stratified 5-fold cross-validation with fixed, pre-registered hyperparameters. Stream A achieved a best AUC-ROC of 0.63 (random forest; permutation p = .049), confirming H1 that acoustic prosody classifies PHQ-8-defined anhedonia above chance at = .05 (uncorrected), though this result does not survive Bonferroni correction for three classifiers (/3 = .017). Bootstrap analysis of {Delta}AUC confirmed non-inferiority relative to Stream B (H2: 95% CI lower bound > -0.10). However, Stream B itself did not achieve above-chance classification (p = .057), so this non-inferiority finding reflects comparable, modest performance across both modalities rather than equivalence to a validated neural biomarker. Pitch variability features (F0) ranked among the top 5 predictors by SHAP value in the gradient-boosted trees model (H3: partially supported). An exploratory combined model (eGeMAPS + recurrence quantification analysis) reached AUC = 0.65 (GBT; p = .032), though this lift was not statistically significant by DeLong test (z = -0.44, p = .66). These results provide initial evidence that vocal prosody, acquired through a standardized, open-source pipeline, carries depression-related information comparable to fMRI-derived ventral striatal activation for binary classification. However, neither streams operationalization isolates anhedonia from general depressive symptomatology, and the cross-modal design compares different constructs across different cohorts. We refer to the classification target throughout as the "anhedonia composite" to acknowledge that PHQ-8 Items 1+2 conflate anhedonia with dysphoria. We discuss these constraints and their implications for the causal framework motivating this work.

16
Evaluating an Adjusting Alcohol Purchase Task as a Brief Measure of Behavioural Economic Demand for Alcohol in a Large Sample of Community Adults

Coelho, S. G.; Belisario, K. L.; Keough, M. T.; MacKillop, J.

2026-07-06 psychiatry and clinical psychology 10.64898/2026.07.02.26357139 medRxiv
Top 0.1%
2.2%
Show abstract

Alcohol demand is commonly assessed using hypothetical alcohol purchase tasks (APTs), from which individual demand curves are constructed and yield multiple indices of reinforcing value. Procedurally, APTs can confer participant burden, and existing brief alternatives cannot produce demand curves or derived indices. Thus, we evaluated a novel, adjusting APT that efficiently and idiographically assesses alcohol demand while preserving the benefits of a full task. Adults reporting past-six-month alcohol use (n=897) completed either the adjusting or full APT, the former utilizing a binary-search-style algorithm to administer six prices from the full APT's price set based on level of alcohol demand. The adjusting APT reduced item burden by 49% and produced well-fitting individual demand curves. Average demand intensity and elasticity estimates did not differ significantly by modality, whereas Omax and breakpoint estimates were significantly higher on the adjusting APT, though only by $3 each. All demand indices from both APTs were positively associated with alcohol use and problems, with similar magnitude by modality. Results provide support for the adjusting APT as a brief measure of alcohol demand that retains demand-curve-based indices of reinforcing value.

17
Predictability and controllability shape aversive learning and stress responses through independent computational mechanisms

Rajput, D.; Felmingham, K.; Sophie Lin, C.-H.; Garrido, M.

2026-09-01 neuroscience 10.64898/2026.08.26.745064 medRxiv
Top 0.1%
2.1%
Show abstract

BACKGROUND: An individual's adaptation to threatening environments under uncertainty is reflected in stress responses. Predictability (the ability to anticipate events) and controllability (the ability to control outcomes) are central to how one adapts, yet their joint influence on aversive learning remains unclear. METHODS: Thirty healthy adults completed a probabilistic aversive learning task in which cue-outcome contingencies varied across levels of predictability and controllability, i.e. whether shock intensity depended on prediction accuracy. Prediction accuracy, reaction time, subjective stress ratings, and skin conductance responses were recorded throughout. Trial-wise learning dynamics were estimated using the Volatile Kalman Filter. RESULTS: Prediction accuracy reduced as environments became less predictable and negatively associated with higher learning rates across predictability levels, with the strongest relationship observed in highly predictable blocks. Skin conductance responses showed that moderately predictable environments elicited responses like those in highly predictable environments when accurate predictions reduced shock intensity, but resembled responses in unpredictable environments when shock intensity was uncontrollable. Model comparison revealed a double dissociation between subjective stress ratings and skin conductance responses. Subjective ratings were best explained by model-derived volatility when prediction accuracy determined shock intensity and by belief uncertainty when it was independent of prediction accuracy, whereas skin conductance responses showed the reverse pattern. Reaction times were best explained by belief uncertainty when predictions influenced shock intensity. Higher anxiety was associated with elevated learning rates in highly and moderately predictable blocks when predictions did not control shock intensity.

18
A double index for assessing symptoms in psychosis based on acoustic and semantic information

Kirdun, M.; He, R.; Demirlek, C.; Garcia-Molina, J. T.; Huppi, R.; Surbeck, W.; Dannecker, N.; Verim, B.; Yalincetin, B.; Ortiz Garcia de la Foz, V.; Ayesa Arriola, R.; Bora, E.; Figueroa-Barra, A. I.; Spaniel, F.; Palaniyappan, L.; Sommer, I. E.; Homan, P.; Hinzen, W.; Palominos, C.

2026-08-12 psychiatry and clinical psychology 10.64898/2026.08.11.26360092 medRxiv
Top 0.1%
1.9%
Show abstract

Recent computational approaches to speech in psychosis generate large multidimensional feature spaces capturing semantic and acoustic aspects of language production. However, the clinical relevance of individual measures often becomes difficult to interpret due to redundancy, interaction effects, and high intercorrelation among features. Building upon previous work constructing a single composite index derived from semantic features based on language model embeddings, we here build a second acoustic index derived from speech acoustic features. Our aim was to evaluate the differential performance of both indices in conjunction in PANSS symptom prediction in psychosis in a cross-linguistic setting, including positive symptoms (P1, P2, P3), negative symptoms (N1, N4, N6), and general measures (G5, G9). The dataset comprised five languages and 221 patients with schizophrenia spectrum disorder (SSD). Both indices showed predictive power for individual PANSS scores, while also demonstrating clinically important complementarity: semantic indices were more strongly associated with positive symptom dimensions (P2, P3, Total Positive), whereas acoustic indices showed stronger relationships with negative and general symptoms (N1, N4, G5, G9). Both domains shared predictive overlap for global measures, such as PANSS Total scores. These findings suggest that both indices capture complementary and partially overlapping dimensions of psychopathology. The proposed composite index framework contributes to the advancement of low-dimensional speech-derived markers of symptom severity variation, potentially informing vulnerability to relapse and remission in psychosis.

19
Antidepressant Maintenance Versus Active Monitoring After Depression Remission: A Decision Analysis Stratified by Relapse Risk and Patient Preferences

Meyerson, W. U.; Cai, T.; Smoller, J. W.

2026-07-20 psychiatry and clinical psychology 10.64898/2026.07.17.26358340 medRxiv
Top 0.1%
1.9%
Show abstract

Importance: Patients who achieve remission from major depressive disorder (MDD) often face a preference-sensitive decision between continued antidepressant maintenance and discontinuation with active monitoring. Quantifying the tradeoff between depression burden and long-term medication exposure may support more individualized shared decision-making. Objective: To quantify tradeoffs between continuous antidepressant maintenance and active monitoring after MDD remission, and to identify preference thresholds favoring each strategy across relapse-risk strata. Design: Individual-level decision-analytic health-state transition model calibrated to randomized maintenance-discontinuation trials and a longitudinal first depressive episode cohort, with a 5-year time horizon. Setting: Outpatient clinical decision after completion of an 8-month continuation phase following remission from MDD. Participants: Adults in remission from MDD, represented across 4 clinically anchored relapse-risk strata ranging from very low risk after a first mild episode to high risk after highly recurrent depression. Exposures: Continuous antidepressant maintenance vs discontinuation with active monitoring and antidepressant restart after detected relapse. Main Outcomes and Measures: Severity-weighted depression-months, antidepressant medication-years, medication-years per depression-month averted, and net benefit across preference thresholds defined as the maximum additional medication-years a patient would be willing to accept to avert 1 depression-month. Results: Continuous maintenance reduced depression burden but required substantially more medication exposure, with efficiency strongly dependent on relapse risk. Medication-years per depression-month averted ranged from 11.8 (95% uncertainty interval [UI], 7.8-19.6) in the very low-risk group to 1.5 (95% UI, 0.8-3.0) in the high-risk group. At a preference threshold of 3 medication-years per depression-month averted, maintenance was preferred for moderate- and high-risk patients; at a threshold of 2, only for high-risk patients; and at a threshold of 1, for no risk group. Conclusions and Relevance: In this decision-analytic model, the value of continuous antidepressant maintenance depended strongly on baseline relapse risk and patient preferences regarding long-term medication exposure. These findings provide a quantitative framework for shared decision-making about antidepressant maintenance after remission from MDD.

20
Prompt Engineering Limitations: Preliminary Evaluation of Large Language Models for Psychotherapy Safety

Ngo, N.; Dao, G.; Sano, A.

2026-07-18 psychiatry and clinical psychology 10.64898/2026.07.16.26358261 medRxiv
Top 0.1%
1.8%
Show abstract

Large Language Models are increasingly used in consumer-facing mental health tools, many of which claim that prompt engineering alone can ensure safe therapeutic behavior. This study evaluates that assumption by testing 20 proprietary and open-source LLMs on high-risk psychiatric scenarios, using prompts grounded in behavioral therapy principles. Prompt engineering reduced some predictable risks, such as explicit endorsement of self-harm, but consistently failed in ambiguous or clinically nuanced situations. Models frequently validated harmful statements, colluded with hallucinations, minimized symptoms, or used stigmatizing language, including in the newest and largest models. These failures reflect structural limitations such as lack of memory, insufficient contextual reasoning, and training-related biases. Prompt engineering alone is therefore insufficient for safe AI-mediated psychotherapy; clinician-guided fine-tuning, integrated safety mechanisms, and system-level oversight will be required. This work provides early evidence motivating deeper clinician-led evaluation and safety-oriented model development.